2026年8月10日星期一

MiniMax AI Review 2026: M2.7 Value King + H3 Video 1, But Don't Rely on It for Coding

MiniMax AI Review 2026: M2.7 Value King + H3 Video 1, But Don't Rely on It for Coding


Last updated: August 11, 2026 | Reading time: 8 minutes




Quick Verdict


| Aspect | Rating |

|--------|--------|

| Price / Value | ⭐⭐⭐⭐⭐ |

| Speed (Output TPS) | ⭐⭐⭐⭐⭐ |

| Agent Tool Calling | ⭐⭐⭐⭐ |

| Video Generation (H3) | ⭐⭐⭐⭐⭐ |

| Coding (Real-world) | ⭐⭐ |

| Complex Reasoning | ⭐⭐ |


Best for: Cost-sensitive startups running large-scale agents, high-concurrency To-C products, video creators/e-commerce teams, and developers who want fully open-source models for self-hosting.


Skip if: You need complex math / rigorous reasoning (research/STEM), top-tier real-world coding, multimodal input, or professional high-quality video production.


---


 What is MiniMax?


MiniMax (parent company: Xiyu Technology) was founded in December 2021 by Junje Yan, former VP of SenseTime. In just 4 years it became one of China's "Big Four" LLM startups — with a distinct profile: overseas-first, cost-obsessed, and full-stack (text + voice + video).


Key numbers that frame its position:


- Listed on the Hong Kong Stock Exchange on Jan 9, 2026 (00100.HK) — "Shanghai's first LLM IPO" and the largest AI-model IPO ever

- IPO price 165 HKD, +109% on day one to 345 HKD, market cap briefly broke 100B HKD

- 1,209× oversubscribed, raised ~5.54B HKD with 14 cornerstone investors (Alibaba, Abu Dhabi Investment Authority)

- Fewer than 400 employees, average age 29, ~$500M total invested since founding — about 1% of what OpenAI spent

- As of Sept 2025: 200M+ personal users, 130K enterprise customers, 200+ countries, overseas revenue >70%

- 2025 revenue $79.04M (+158.9%), but cumulative losses of ~$1.32B — "high growth, high losses"


Product map: Talkie (domestic: 星野) for AI companion social, Hailuo AI for consumer video/creation, and models from M1 up to M2.7 and H3. Its Speech 02 voice model once topped the global charts.


---


 Key Features (2026)


 1. Full Model Family: M2.5 Open / M2.7 Flagship / H3 Video


| Model | Released | Focus | Key Numbers |

|-------|----------|-------|-------------|

| MiniMax M1 | May 2025 | Reasoning agent | 2 globally on MRCR / LongBench |

| MiniMax M2 | Oct 2025 | Open-source leader | Top-5 Artificial Analysis, 1 open-source |

| M2.5 | Feb 2026 | Open-source flagship | Apache 2.0, 256B (45.9B active), SWE-bench 80.2% |

| M2.7 | Mar 2026 | Closed flagship | First "self-evolving" model, 1/50 price |

| H3 (Hailuo 3.0) | Jul 2026 | Video generation | Native 2K, 1 global video editing |


 2. M2.5: The "Production-Grade Native Agent" Model


- Fully open source (Apache 2.0), a complete 256B-parameter / 45.9B-active MoE — not a gutted "open-weights" release; deployable via vLLM / SGLang

- SWE-bench Verified 80.2%, Multi-SWE-bench 51.3% — beats Claude Opus 4.6

- Positioned as a production-grade Agent model: within a day of release, users built 10,000+ expert agents on the MiniMax Agent platform

- Lightning version hits 100+ TPS output (~2× mainstream): $10,000 can theoretically power 4 agents working non-stop for a year


 3. M2.7: The First "Self-Evolving" Commercial LLM


- Used its internal OpenClaw framework to self-optimize over 100+ rounds (rewriting scaffolding code, tuning sampling params, building its own operational rules) — +30% on internal evals, zero human intervention. "AI training itself."

- Built for Agent tool-calling and engineering coding: SWE-bench 78% (Opus 4.6: 55%), Terminal-Bench 2 at 82.4%, tool-calling accuracy 75.8%

- ~100 token/s (~2× mainstream), only ~10B active parameters

- Caveat: text-only (no multimodal input), and systematically weak at complex math / multi-step reasoning (HLE just 28)


 4. H3 (Hailuo 3.0): Video Generation, 1 Global Editing


Announced at WAIC on July 31, 2026; open-sourced Aug 3:

- Native 2K, 24fps, 5–15s clips per generation (extendable to ~30s), 21:9 to 9:16 aspect ratios

- Omni-Reference: up to 9 reference images + 3 videos + 3 audio clips at once to lock style, character, motion, and voice consistency

- Single-pass synchronized audio: dialogue + SFX + ambience in one generation — no separate audio pipeline

- 0.8 yuan/second (2K) — about one-third the cost of rival flagship video models

- 1 in global video editing on Artificial Analysis; Arena blind-test image-to-video 1476 (just 2 points behind leader Seedance 2.0's 1478); 1 among all open models (~280 points above the 2 open model)


 5. Hailuo AI: Consumer Creation App / Web


- Text creation + AI image + AI video generation, plus signature voice cloning / timbre imitation

- Free tier: ~10 generations/day, no watermark, usable without registering; each clip generated in ~30s–5min — far faster than rivals (no 6-hour queues)

- Excellent physics simulation (liquid/fabric/texture), rich style templates (cyberpunk, ink-wash landscape), strong Chinese-language handling, zero learning curve


---


 Pricing (2026)


 API per 1M tokens


| Model | Input | Output | Notes |

|-------|-------|--------|-------|

| M2.7 Standard | $0.30 | $1.20 | Cache $0.06; ~1/50 of Claude Opus 4.6 |

| M2.7 Highspeed | $0.60 | $2.40 | Same ability, 2× price, faster |

| M2.5 Official | ~$0.30 | ~$2.40 | Lightning: 100+ TPS |

| M2.5 Alibaba Cloud | 2.1 yuan | 8.4 yuan | Cache-hit 0.21–0.42 yuan |


> Cost math: 200 code-review tasks per month costs ~$1.62 on M2.7 vs ~$94.50 on Claude Opus 4.6 — nearly 60× difference. M2.7 input is ~1/17 of GPT-5.5 and 1/50 of Claude Opus 4.6.


[IMAGE-1: MiniMax M2.7 API pricing comparison table vs Claude / GPT per million tokens]


 Hailuo AI Subscription


- Free: ~10 gens/day, 720p, 6s clips (Pro 10s), no watermark

- International: ~$9.99/month

- Premium (尊享): ~1,399 yuan/month (20,000 credits/mo, unlimited Hailuo full model family)

- ⚠️ Users have complained the top-tier annual plan's lip-sync feature is severely defective and generation is very slow — false-advertising allegations


---


 What We Liked (Test Results)


 ✅ Price Slasher: 1/50 of Claude's Price


M2.7 input at $0.30/M is 1/50 of Claude Opus 4.6 and 1/17 of GPT-5.5; ~64% cheaper than GLM-5 overall. Crushing advantage for cost-sensitive Agent and To-C products.


 ✅ Fastest Output: ~2× Mainstream


M2.7 at ~100 token/s and M2.5 Lightning at 100+ TPS are about twice the mainstream speed — a visible difference in high-concurrency, real-time interaction.


 ✅ Strong Agent Tool Calling


Terminal-Bench 2 at 82.4% (Opus: 74.1%), tool-calling accuracy 75.8%, clearly ahead in system-level/DevOps operations. SWE-bench Pro 56.22% is on par with Opus.


 ✅ 1 Open-Source Video Generation


H3 ranks 1 globally in video editing on Artificial Analysis, dominating open models (~280 pts clear), at 0.8 yuan/second with native 2K and single-pass audio.


 ✅ Groundbreaking "Self-Evolution"


M2.7 is the first commercial model to train itself — +30% internal evals with zero human intervention. A genuinely imaginative technical direction.


 ✅ Proven Overseas Product Ability


>70% of revenue is overseas; Talkie and the 1-global Speech 02 voice model prove MiniMax's consumer products survive the global market.


 ✅ Deep Voice/Multimodal Foundation


Speech 02 once beat OpenAI to 1 globally; Hailuo 02 video ranked top-3. Voice cloning / timbre replication is mature at consumer level.


---


 What We Didn't Like


 ❌ Complex Reasoning Is a Hard Weakness


HLE score of just 28 — systematically weak at complex math and multi-step logic. Not suitable for research / STEM / rigorous reasoning. This is its biggest gap versus DeepSeek, GPT, and Claude.


 ❌ Coding "Scores High, Fails in the Wild"


- Benchmarks look great (SWE-bench 78%), but in a real-world GitHub project shootout it scored only 76 (B-grade) — the backend test failed to compile, while DeepSeek V4-Pro and GLM-5.1 both scored S-grade (90+)

- In a security-code analysis test, M2.7 burned the most tokens (1.09M — 5× GLM-5.1) yet found zero of the core vulnerabilities — direction of reasoning matters more than volume


 ❌ M2.7 Is Text-Only


The flagship lacks image/video input — no native multimodal, which clashes with its "all-rounder" marketing.


 ❌ Hailuo Free Tier Is Limited


6-second clips can't do narrative, 720p looks grainy when upscaled, complex motion breaks down (limb contortion, mangled fingers), no camera-control, 16:9 only, and peak-hour queues. Overall review score 7.4/10 — ranked 6th of 8 free AI video tools.


 ❌ Commercialization Pressure


Cumulative losses of ~$1.32B; 2025 compute costs ~1B yuan; the company admits net losses will keep growing. Whether "high growth + high losses" is sustainable is the biggest question mark.


 ❌ Membership Controversy


Hailuo's premium annual plan (~1,399 yuan/month) drew complaints about defective lip-sync, very slow generation, and false-advertising allegations.


 ❌ Inconsistent Spec Info


M2.7 context window is reported as 200K / 262K / 1M; release date as Mar or Apr 2026. Official communication lacks consistency.


---


 MiniMax vs DeepSeek vs GLM vs Claude (2026)


| Dimension | MiniMax M2.7 | GLM-5.1 | DeepSeek V4-Pro | Claude Opus 4.6 |

|-----------|--------------|---------|-----------------|-----------------|

| Input / 1M tokens | $0.30 (lowest) | ~$1.74 | $1.74 (Flash tier cheaper) | $15 |

| Output speed | ~100 TPS (fastest) | Slow | Medium | Medium |

| Active params | ~10B (smallest) | 40B | 49B | — |

| Context | ~200K | 200K | 1M (longest) | 200K |

| SWE-bench Pro | 56.2% | 58.4% (1 open) | 55.4% | ~55% |

| Real-world GitHub | 76 (B, backend failed) | 90 (S) | 92 (S) | — |

| Complex math / reasoning | ❌ Weak (HLE 28) | Strong | Excellent | Excellent |

| Video generation | H3 1 global | ❌ | ❌ | ❌ |

| Open source | ✅ M2.5 open | ✅ MIT | ✅ MIT | ❌ |


How to choose:


- Cost-sensitive / high-concurrency / real-time → MiniMax M2.7 (1/50 price, 2× speed)

- Video creation / e-commerce ads → MiniMax H3 (1 open, 0.8 yuan/s)

- Mainline coding / engineering reliability → GLM-5.1 or DeepSeek V4 (S-grade in the wild)

- Research / rigorous reasoning → DeepSeek V4 / Claude (crushing math advantage)


---


 Who Should Use MiniMax?


 ✅ Good fit


- Cost-sensitive startup teams: large-scale agent loops and batch jobs — M2.7 cuts cost by an order of magnitude

- High-concurrency To-C products: real-time support, conversational UI — 2× speed is a visible UX win

- Video creators / e-commerce operators: H3 is open-source and deployable locally; 0.8 yuan/s batch output with strong ad/product-shots

- Self-hosting developers: M2.5 and H3 fully open, vLLM/SGLang local deployment, no hardware lock-in

- Voice-cloning / timbre apps: Speech series is 1 globally with mature consumer UX


 ❌ Poor fit


- Researchers / STEM professionals: complex math and multi-step reasoning are hard weaknesses

- Mainline developers: high scores but backend fails in the wild; inefficient reasoning — don't make it your primary coding model

- Multimodal input needs: M2.7 is text-only; image/video understanding requires a separate path

- High-quality video pros: free Hailuo is 6s/720p; professional work needs premium or rival tools like Kling

- Conservatives / long-term planners: losses keep widening; the business model isn't proven yet


---


 Final Verdict: Is MiniMax Worth It?


Price slasher + video king — but don't rely on it for coding or reasoning.

One line: M2.7 handles "more, faster, cheaper"; H3 handles "1 global video"; let DeepSeek/GLM cover the weak spots.


- If you're a developer or founder, M2.7's 1/50 price + 2× speed + strong Agent make it the value king for large-scale agents and To-C products

- If you're a video creator / e-commerce seller, H3's 1 open source + 0.8 yuan/s + native 2K is a commercial production weapon

- If you do research or rigorous reasoning, it's not for you — complex math is a hard stop

- If coding is your main job, use M2.7 as a helper and DeepSeek/GLM as the main engine


| Use Case | Recommendation |

|----------|----------------|

| Cost-efficient agent scaling | ✅ Highly recommended |

| Real-time To-C chat | ✅ Highly recommended |

| Video / ad batch production | ✅ Highly recommended |

| Mainline coding | ⚠️ Use GLM/DeepSeek instead |

| Research / STEM reasoning | ⚠️ Use DeepSeek/Claude |

| Consumer video (free tier) | ✅ Good for light use |


MiniMax's core strengths: extreme value, fastest speed, 1 global video, full-stack open source. Its weaknesses: weak complex reasoning, real-world coding misses, text-only flagship, mounting losses. It's the "king of cost and speed," not the "ceiling of capability" — use it in the right scenario and it's a great tool.


Scorecard: Overall 8.2/10 · Price/Value 9.5 · Speed 9.5 · Video (H3) 9.5 · Agent Tool-Calling 8.5 · Real-world Coding 5.5 · Complex Reasoning 4.5


---


 FAQ


Q: Is MiniMax a public company?

A: Yes. Xiyu Technology listed on the HKEX on Jan 9, 2026 (00100.HK), +109% on day one with market cap briefly above 100B HKD — the largest AI-model IPO in history.


Q: How much does MiniMax M2.7 cost?

A: Standard: $0.30/M input, $1.20/M output, $0.06 cache. Highspeed: $0.60/$2.40. Input is about 1/50 of Claude Opus 4.6.


Q: Is MiniMax open source?

A: M2.5 (Apache 2.0) and H3 are fully open source and self-hostable. M2.7 is a closed-source flagship (some channels report MIT).


Q: Which is better for coding — MiniMax, DeepSeek, or GLM?

A: SWE-bench Pro is close across all three (GLM-5.1 58.4% leads), but in a real-world GitHub shootout DeepSeek V4-Pro (92) and GLM-5.1 (90) were S-grade while MiniMax scored 76 (B, backend failed). For mainline coding choose DeepSeek or GLM.


Q: How good is MiniMax's H3 video model?

A: H3 (Hailuo 3.0) is 1 globally in video editing on Artificial Analysis and 1 among open models by a huge margin — native 2K/24fps, single-pass synchronized audio, 0.8 yuan/s (1/3 of rivals). Open-sourced Aug 3, 2026.


Q: Is Hailuo AI free?

A: The free tier gives ~10 generations/day, 720p, 6s clips, no watermark — enough for light use. Premium is ~$9.99/month; the top tier runs ~1,399 yuan/month for unlimited full-model access.


Q: Is MiniMax good for research/reasoning?

A: No. M2.7 is systematically weak at complex math and multi-step logic (HLE just 28). For research/STEM use DeepSeek or Claude.


---


This review was last updated on August 11, 2026. Product features and pricing are subject to change. Always check the official website for the latest information.


This post is part of our AI Tools Review Series. More reviews coming: DeepSeek, Doubao, and more.


---


Sources:

- [Sina Finance: MiniMax M2.5 release — on par with Claude Opus 4.6 at ~$0.3/M input](https://finance.sina.com.cn/tech/2026-02-13/doc-inhmrnzr3313137.shtml)

- [Xinhua Economic Daily: MiniMax M2.5 globally open-sourced with local deployment](http://jjckb.xinhuanet.com/20260213/f083116f8a184cb4919d6b783df219ea/c.html)

- [Tencent Cloud Dev Community: MiniMax M2.7 API guide & cost analysis (2026)](https://cloud.tencent.cn/developer/article/2656054)

- [ofox.ai: MiniMax M2.7 self-evolving model deep dive](https://ofox.ai/zh/blog/minimax-m2-7-self-evolving-ai-model-2026/)

- [Eastmoney: MiniMax third-gen video model open-sourced, video editing 1 globally](https://finance.eastmoney.com/a/202608033829888273.html)

- [36Kr AI Review: Hailuo AI hands-on test (2026)](https://36aidianping.com/note-detail/3568010593721428)

- [Tencent Cloud / Zhihu: Four-way shootout — DeepSeek V4-Pro / GLM-5.1 / MiniMax M2.7](https://cloud.tencent.com.cn/developer/article/2660044)

- [GitHub LLM-Comparison: real-world coding shootout data](https://github.com/Sj295/LLM-Comparison)


没有评论:

发表评论

Microsoft MAI Review 2026: The 7-Model Family That Says Goodbye to OpenAI — Trillion-Parameter Flagship, Now Default in GitHub Copilot

Microsoft MAI Review 2026: The 7-Model Family That Says Goodbye to OpenAI — Trillion-Parameter Flagship, Now Default in GitHub ...